Papers with macro F1 metrics
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild (2022.lrec-1)
Copied to clipboard
| Challenge: | GINCO is a new training dataset for automatic genre identification based on 1,125 crawled Slovenian web documents that consist of 650,000 words. |
| Approach: | They propose to use 1,125 crawled Slovenian web documents to train a new genre classification system based on a GINCO training dataset . |
| Outcome: | The proposed classifiers perform better on the 1,125 crawled Slovenian web documents than the existing models and achieve higher scores on the task. |